Tag: ai-alignment

Blog Posts

Why Alignment Verification Might Be Fundamentally Broken

Turing proved in 1936 that universal verification is impossible. Now we're trying it anyway, on AI systems that adapt to whatever detection we point at them.

Hand me a detector f and I can build a program g that defeats it. The same trap catches alignment testing: every test you run is one more signal telling the model humans are watching.

The Yard, The Sparkly Hat, and The Doomsday Clock

Most AI doom talk comes from people with money on it. The industry titans hype their own power, and a handful of obscure nonprofits forecast the apocalypse to keep the donations rolling in. What caught my attention were three writers standing outside both rackets.

Freddie deBoer plays the skeptic and mocks the whole thing with his "Shitting-in-the-Yard Challenge." Scott Alexander, a rationalist, takes MIRI's doomsday math and turns it into a toddler behind the wheel of a Ferrari. Daniel Kokotajlo walked away from millions in OpenAI equity to warn about a 2027 AGI arms race.

None of them agree on what's coming, and they'd probably argue about it for hours. But they land on the same worry: our institutions and incentives aren't ready for what we're building. When three people with nothing to gain point at the same spot, I pay attention, even if they can't tell me exactly what's wrong.